Tag
3 articles
OpenAI and Broadcom have introduced Jalapeño, a custom AI chip optimized for large language model inference, promising enhanced performance and efficiency. The chip represents a significant advancement in AI hardware infrastructure.
OpenAI and Broadcom have unveiled the 'Jalapeño' chip, a custom processor designed to accelerate large language model inference, with deployment expected by late 2026.
Researchers at UC San Diego introduce DFlash, a new speculative decoding technique that drafts whole token blocks in parallel, achieving up to 15x throughput improvement on NVIDIA Blackwell.